Papers by Andrew D Wilson
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames (2025.emnlp-main)
Copied to clipboard
Sahithya Ravi, Gabriel Herbert Sarch, Vibhav Vineet, Andrew D Wilson, Balasaravanan Thoravi Kumaravel
| Challenge: | Disjoint-3DQA evaluates the spatial reasoning ability of embodied AI assistants based on egocentric video . it aims to catalyze future research at the intersection of vision, language, and embodie . |
| Approach: | They propose a generative QA benchmark that evaluates the ability of embodied AI assistants to integrate spatial cues across time by asking object pairs that are not co-visible in the same frame. |
| Outcome: | The proposed benchmark compares seven state-of-the-art VLMs and finds that they lag behind human performance by 28%, with steeper declines as the temporal gap widens. |
Grounding Task Assistance with Multimodal Cues from a Single Demonstration (2025.findings-acl)
Copied to clipboard
Gabriel Herbert Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, Andrew D Wilson
| Challenge: | RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior. |
| Approach: | They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests. |
| Outcome: | The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering. |